JAMIA Open
◐ Oxford University Press (OUP)
All preprints, ranked by how well they match JAMIA Open's content profile, based on 42 papers previously published here. The average preprint has a 0.07% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Gilson, A.; Schulz, W. L.; Lopez, K.; Young, P.; Pandya, S.; Coppi, A.; Chartash, D.; Fiellin, D.; D'Onofrio, G.; Taylor, R. A.
Show abstract
ObjectiveWe aimed to discover computationally-derived phenotypes of opioid-related patient presentations to the emergency department (ED) via clinical notes and structured electronic health record (EHR) data. MethodsThis was a retrospective study of ED visits from 2013-2020 across ten sites within a regional healthcare network. We derived phenotypes from visits for patients 18 years of age with at least one prior or current documentation of an opioid-related diagnosis. Natural language processing was used to extract clinical entities from notes, which were combined with structured data within the EHR to create a set of features. We performed Latent Dirichlet allocation to identify topics within these features. Groups of patient presentations with similar attributes were identified by cluster analysis. ResultsIn total 82,577 ED visits met inclusion criteria. The 30 topics discovered ranged from those related to substance use disorder, chronic conditions, mental health, and medical management. Clustering on these topics identified nine unique cohorts with one-year survivals ranging from 84.2-96.8%, rates of one-year ED returns from 9-34%, rates of one-year opioid event 10-17%, rates of medications for opioid use disorder from 17-43%, and a median Carlson comorbidity index of 2-8. Two cohorts of phenotypes were identified related to chronic substance use disorder, or acute overdose. ConclusionsOur results indicate distinct phenotypic clusters with varying patient-oriented outcomes which provide future targets better allocation of resources and therapeutics. This highlights the heterogeneity of the overall population, and the need to develop targeted interventions for each population.
Pfaff, E.; Bradford, R.; Clark, M.; Balhoff, J. P.; Wang, R.; Preisser, J. S.; Walters, K.; Nielsen, M. E.
Show abstract
BackgroundComputable phenotypes are increasingly important tools for patient cohort identification. As part of a study of risk of chronic opioid use after surgery, we used a Resource Description Framework (RDF) triplestore as our computable phenotyping platform, hypothesizing that the unique affordances of triplestores may aid in making complex computable phenotypes more interoperable and reproducible than traditional relational database queries. To identify and model risk for new chronic opioid users post-surgery, we loaded several heterogeneous data sources into a Blazegraph triplestore: (1) electronic health record data; (2) claims data; (3) American Community Survey data; and (4) Centers for Disease Control Social Vulnerability Index, opioid prescription rate, and drug poisoning rate data. We then ran a series of queries to execute each of the rules in our "new chronic opioid user" phenotype definition to ultimately arrive at our qualifying cohort. ResultsOf the 4,163 patients in the denominator, our computable phenotype identified 248 patients as new chronic opioid users after their index surgical procedure. After validation against charts, 228 of the 248 were revealed to be true positive cases, giving our phenotype a PPV of 0.92. ConclusionWe successfully used the triplestore to execute the new chronic opioid user phenotype logic, and in doing so noted some advantages of the triplestore in terms of schemalessness, interoperability, and reproducibility. Future work will use the triplestore to create the planned risk model and leverage the additional links with ontologies, and ontological reasoning.
Honerlaw, J.; Ho, Y.-L.; Fontin, F.; Gosian, J.; Maripuri, M.; Murray, M.; Sangar, R.; Galloway, A.; Zimolzak, A. J.; Whitbourne, S. B.; Casas, J. P.; Ramoni, R.; Gagnon, D. R.; Cai, T.; Liao, K. P.; Gaziano, J. M.; Muralidhar, S.; Cho, K.
Show abstract
The development of phenotypes using electronic health records is a resource intensive process. Therefore, the cataloging of phenotype algorithm metadata for reuse is critical to accelerate clinical research. The Department of Veterans Affairs Office of Research and Development has developed a phenomics knowledgebase library, CIPHER (Centralized Interactive Phenomics Research), which improves upon existing phenomics library models to help advance innovation in clinical research by using the CIPHER phenotype collection standard. The CIPHER standard was iteratively developed with phenomics experts and has been used to capture over 5,000 phenotypes. We describe the development of the CIPHER standard for phenotype metadata collection, its current application to the largest healthcare system in the United States, and the future expansion of the CIPHER knowledgebase as a public resource for phenotyping.
Elhussein, A.; Hripcsak, G.
Show abstract
The dimensionality of electronic health record (EHR) data continues to grow as more clinical variables are recorded, often resulting in redundancy, sparsity, and analytical intractability. In this study, we apply non-negative matrix factorization (NMF) to a high-dimensional laboratory dataset of patients with type II diabetes to estimate the minimum latent dimensionality required to preserve clinically meaningful information. Using both within-patient imputation and across-patient generalization tasks, we evaluate the ability of the learned representations to reconstruct two key clinical lab values: blood glucose and HbA1c. Our findings show that clinically acceptable accuracy can be achieved with a dimensionality reduction of up to 80% and a dimensionality of 230 to 300, supporting the presence of a compact, low-dimensional latent structure underlying high-dimensional clinical data.
Waters, R.; Malecki, S.; Lail, S.; Mak, D.; Saha, S.; Jung, H. Y.; Razak, F.; Verma, A.
Show abstract
ObjectivePatient data repositories often assemble medication data from multiple sources, necessitating standardization prior to analysis. We implemented and evaluated a medication standardization procedure for use with a wide range of pharmacy data inputs across all drug categories, which supports research queries at multiple levels of granularity. MethodsThe GEMINI-RxNorm system automates the use of multiple RxNorm tools in tandem with other datasets to identify drug concepts from pharmacy orders. GEMINI-RxNorm was used to process 2,090,155 pharmacy orders from 245,258 hospitalizations between 2010 and 2017 at 7 hospitals in Ontario, Canada. The GEMINI-RxNorm system matches drug-identifying information from pharmacy data (including free-text fields) to RxNorm concept identifiers. A user interface allows researchers to search for drug terms and returns the relevant original pharmacy data through the matched RxNorm concepts. Users can then manually validate the predicted matches and discard false positives. We designed the system to maximize recall (sensitivity) and enable excellent precision (positive predictive value) with minimal manual validation. We compared the performance of this system to manual coding (by a physician and pharmacist) of 13 medication classes. ResultsManual coding was performed for 1,948,817 pharmacy orders and GEMINI-RxNorm successfully returned 1,941,389 (99.6%) orders. Recall was greater than 98.5% in all 13 drug classes, and the F-Measure and precision remained above 90.0% in all drug classes, facilitating efficient manual review to achieve 100.0% precision. GEMINI-RxNorm saved time substantially compared to manual standardization, reducing the time taken to review a pharmacy order row from an estimated 30 seconds to 5 seconds and reducing the number of rows needed to be reviewed by up to 99.99%. Discussion and ConclusionGEMINI-RxNorm presents a novel combination of RxNorm tools and other datasets to enable accurate, efficient, flexible, and scalable standardization of pharmacy data. By facilitating efficient minimal manual validation, the GEMINI-RxNorm system can allow researchers to achieve near-perfect accuracy in medication data standardization.
Dasgupta, R.
Show abstract
BackgroundTraditional pharmacovigilance methods based on biostatistical approaches systematically exclude outliers and rare events, potentially missing critical safety signals. These methods fail to detect micro-clusters of adverse events and comorbidity patterns that may indicate serious but low-frequency adverse drug reactions (ADRs). We introduce the concept of absurdity signal detection - the identification of statistically anomalous but clinically significant adverse event patterns that conventional methods dismiss as outliers. MethodsWe developed an ensemble machine learning framework combining five distinct algorithms (Random Forest, Gradient Boosting, XGBoost, Neural Networks, and Support Vector Machines) to analyze FDA Adverse Event Reporting System (FAERS) data. The system employs outlier-inclusive modeling, multi-dimensional cluster detection, and severity-weighted propensity scoring. We validated our approach on Losartan, analyzing 500 adverse event reports to detect absurdity signals that may have been missed by conventional biostatistical surveillance. ResultsOur ensemble approach achieved 75% accuracy in identifying high-risk adverse events, with the best-performing model successfully detecting 15 distinct absurdity signals. The top five identified events were: cough (propensity score 1.525), angioedema (1.298), insomnia (1.290), nausea (1.180), and hyperkalemia (1.114). Notably, our method identified several rare but severe ADRs that would have been excluded as statistical outliers in traditional disproportionality analyses. The ensemble approach demonstrated superior performance compared to individual models, with inter-model agreement providing an additional confidence metric for signal validation. ConclusionsMachine learning-based absurdity signal detection offers a paradigm shift in pharmacovigilance by preserving and analyzing rare adverse events rather than excluding them. This approach has significant implications for patient safety, potentially preventing serious adverse events in vulnerable populations with atypical response profiles. Our methodology is scalable, validated against FDA data sources, and provides a framework for real-time safety monitoring in the $138 billion pharmaceutical industry. Future work will extend this approach to drug-drug interaction detection and personalized risk stratification.
Wasz, M.; Shankar, P. R. V.; Sprouse, E.; Kirchner, L.; Garber, M.; Jones, J.; Mandl, K.; McMurry, A.
Show abstract
ObjectiveDevelop a near-comprehensive opioid medications valueset for population measures of opioid related treatments and outcomes. The opioid valueset should be free, open source, and conform to the RxNorm standard federally mandated in every US-certified electronic health record. Materials and MethodsCumulus opioid valueset was manually curated by the authors and expanded using computer assisted curation. Opioid classifier rules were developed to select opioid RxNorm concepts with known opioid receptor interactions, ingredients, keywords, and drug product formulations. Twelve publicly available valuesets were used to develop and validate the Cumulus opioid valueset. Validation accuracy was measured against a corpus of opioid medication orders and non-opioid pain relievers. ResultsCumulus opioid valueset recall was >99.9% when measured against opioid prescription RxNorm codes from UC Davis Health and Brigham and Womens Hospital. Cumulus opioid valueset was 100% specific compared to three valuesets of non-opioid pain relievers. Discussion and ConclusionTo the authors knowledge, Cumulus opioid valueset is the largest publicly available valueset of opioid medications (8,926 RxNorm concepts). The intended use of this opioid valueset is for population health measures of opioid medications and related patient outcomes.
Shahid, U.; Parde, N.; Smith, D. L.; Dickinson, G.; Bianco, J.; Thorpe, D.; Hota, M.; Afshar, M.; Karnik, N. S.; chhabra, n.
Show abstract
ObjectivesThe accurate identification of Emergency Department (ED) encounters involving opioid misuse is critical for health services, research, and surveillance. We sought to develop natural language processing (NLP)-based models for the detection of ED encounters involving opioid misuse. MethodsA sample of ED encounters enriched for opioid misuse was manually annotated and clinical notes extracted. We evaluated classic machine learning (ML) methods, fine-tuning of publicly available pretrained language models, and a previously developed convolutional neural network opioid classifier for use on hospitalized patients (SMART-AI). Performance was compared to ICD-10-CM codes. Both raw text and text transformed to the United Medical Language System were evaluated. Face validity was evaluated by term feature importance. ResultsThere were 1123 encounters used for training, validation, and testing. Of the classic ML methods, XGBoost had the highest AU_PRC (0.936), accuracy (0.887), and F1 score (0.863) which outperformed ICD-10-CM codes [accuracy 0.870; F1 0.830]. Logistic regression, support vector machine, and XGBoost models had higher AU_PRC using transformed text, while decision trees performed better using raw text. Excluding XGBoost, fine-tuned pre-trained language models outperformed classic ML methods. The best performing model was the fine-tuned SMART-AI based model with domain adaptation [AU_PRC 0.948; accuracy 0.882; F1 0.851]. Explainability analyses showed the most predictive terms were heroin, opioids, alcoholic intoxication, chronic, cocaine, opiates, and suboxone. ConclusionsNLP-based models outperform entry of ICD-10-CM diagnosis codes for the detection of ED encounters with opioid misuse. Fine tuning with domain adaptation for pre-trained language models resulted in improved performance.
Guo, Y.; He, X.; Lyu, T.; Zhang, H.; Wu, Y.; Yang, X.; Chen, Z.; Markham, M. J.; Modave, F.; Xie, M.; Hogan, W.; Harle, C. A.; Shenkman, E. A.; Bian, J.
Show abstract
Transgender and gender nonconforming (TGNC) individuals face significant marginalization, stigma, and discrimination. Under-reporting of TGNC individuals is common since they are often unwilling to self-identify. Meanwhile, the rapid adoption of electronic health record (EHR) systems has made large-scale, longitudinal real-world clinical data available to research and provided a unique opportunity to identify TGNC individuals using their EHRs, contributing to a promising routine health surveillance approach. Built upon existing work, we developed and validated a computable phenotype (CP) algorithm for identifying TGNC individuals and their natal sex (i.e., male-to-female or female-to-male) using both structured EHR data and unstructured clinical notes. Our CP algorithm achieved a 0.955 F1-score on the training data and a perfect F1-score on the independent testing data. Consistent with the literature, we observed an increasing percentage of TGNC individuals and a disproportionate burden of adverse health outcomes, especially sexually transmitted infections and mental health distress, in this population.
Bhasuran, B.; Schmolly, K.; Kapoor, Y.; Jayakumar, N.; Doan, R.; Amin, J.; Meninger, S.; Cheng, N.; Deering, R.; Anderson, K.; Beaven, S.; Wang, B.; Rudrapatna, V. A.
Show abstract
ImportanceAcute Hepatic Porphyria (AHP) is a group of rare but treatable conditions associated with diagnostic delays of fifteen years on average. The advent of electronic health records (EHR) data and machine learning (ML) may improve the timely recognition of rare diseases like AHP. However, prediction models can be difficult to train given the limited case numbers, unstructured EHR data, and selection biases intrinsic to healthcare delivery. ObjectiveTo train and characterize models for identifying patients with AHP. Design, Setting, and ParticipantsThis diagnostic study used structured and notes-based EHR data from two centers at the University of California, UCSF (2012-2022) and UCLA (2019-2022). The data were split into two cohorts (referral, diagnosis) and used to develop models that predict: 1) who will be referred for testing of acute porphyria, amongst those who presented with abdominal pain (a cardinal symptom of AHP), and 2) who will test positive, amongst those referred. The referral cohort consisted of 747 patients referred for testing and 99,849 contemporaneous patients who were not. The diagnosis cohort consisted of 72 confirmed AHP cases and 347 patients who tested negative. Cases were female predominant and 6-75 years old at the time of diagnosis. Candidate models used a range of architectures. Feature selection was semi-automated and incorporated publicly available data from knowledge graphs. Main Outcomes and MeasuresF-score on an outcome-stratified test set ResultsThe best center-specific referral models achieved an F-score of 86-91%. The best diagnosis model achieved an F-score of 92%. To further test our model, we contacted 372 current patients who lack an AHP diagnosis but were predicted by our models as potentially having it ([≥] 10% probability of referral, [≥] 50% of testing positive). However, we were only able to recruit 10 of these patients for biochemical testing, all of whom were negative. Nonetheless, post hoc evaluations suggested that these models could identify 71% of cases earlier than their diagnosis date, saving 1.2 years. Conclusions and RelevanceML can reduce diagnostic delays in AHP and other rare diseases. Robust recruitment strategies and multicenter coordination will be needed to validate these models before they can be deployed. KEY POINTSO_ST_ABSQuestionC_ST_ABSCan machine learning help identify undiagnosed patients with Acute Hepatic Porphyria (AHP), a group of rare diseases? FindingsUsing electronic health records (EHR) data from two centers we developed models to predict: 1) who will be referred for AHP testing, and 2) who will test positive. The best models achieved 89-93% accuracy on the test set. These models appeared capable of recognizing 71% of the cases earlier than their true diagnosis date, reducing diagnostic delays by an average of 1.2 years. MeaningMachine learning models trained using EHR data can help reduce diagnostic delays in rare diseases like AHP.
Essaid, S.; Andre, J.; Brooks, I. M.; Hohman, K. H.; Hull, M.; Jackson, S. L.; Kahn, M. G.; Kraus, E. M.; Mandadi, N.; Martinez, A. K.; Mui, J. Y.; Zambarano, B.; Soares, A.
Show abstract
ObjectiveThe Multi-State EHR-Based Network for Disease Surveillance (MENDS) is a population-based chronic disease surveillance distributed data network that uses institution-specific extraction-transformation-load (ETL) routines. MENDS-on-FHIR examined using Health Language Sevens Fast Healthcare Interoperability Resources (HL7(R) FHIR(R)) and US Core Implementation Guide (US Core IG) compliant resources derived from the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM) to create a standards-based ETL pipeline. Materials and MethodsThe input data source was a research data warehouse containing clinical and administrative data in OMOP CDM Version 5.3 format. OMOP-to-FHIR transformations, using a unique JavaScript Object Notation (JSON)-to-JSON transformation language called Whistle, created FHIR R4 V4.0.1/US Core IG V4.0.0 conformant resources that were stored in a local FHIR server. A REST-based Bulk FHIR $export request extracted FHIR resources to populate a local MENDS database. ResultsEleven OMOP tables were used to create 10 FHIR/US Core compliant resource types. A total of 1.13 trillion resources were extracted and inserted into the MENDS repository. A very low rate of non-compliant resources was observed. DiscussionOMOP-to-FHIR transformation results passed validation with less than a 1% non-compliance rate. These standards-compliant FHIR resources provided standardized data elements required by the MENDS surveillance use case. The Bulk FHIR application programming interface (API) enabled population-level data exchange using interoperable FHIR resources. The OMOP-to-FHIR transformation pipeline creates a FHIR interface for accessing OMOP data. ConclusionMENDS-on-FHIR successfully replaced custom ETL with standards-based interoperable FHIR resources using Bulk FHIR. The OMOP-to-FHIR transformations provide an alternative mechanism for sharing OMOP data. LAY ABSTRACTMany chronic conditions, such as hypertension, obesity, and diabetes are becoming more prevalent, especially in high-risk individuals, such as minorities and low-income patients. Public health surveillance networks measure the presence of specific conditions repeatedly over time, seeking to detect changes in the amount of a disease conditions so that public health officials can implement new early-prevention programs or evaluate the impact of an existing prevention program. Data stored in electronic health records (EHRs) could be used to measure the presence of health conditions, but significant technical barriers make current methods for data extraction laborious and costly. HL7 BULK FHIR is a new data standard that is required to be available in all commercial EHR systems in the United States. We examined the use of BULK FHIR to provide EHR data to an existing public health surveillance network called MENDS. We found that HL7 BULK FHIR can provide the necessary data elements for MENDS in a standardized format. Using HL7 BULK FHIR could significantly reduce barriers to data for public health surveillance needs, enabling public health officials to expand the diversity of locations and patient populations being monitored.
Xu, J.; Hwang, Y. M.; Dormoy, I.; Jing, S. L.; Pillai, M.; Curtin, C. M.; Hernandez-Boussard, T.
Show abstract
Ensuring fairness across diverse patient populations is a fundamental challenge for clinical AI systems, yet current fairness evaluation approaches create critical blind spots. Group-level metrics capture systemic disparities but miss patient-level variations, while individual fairness frameworks ensure consistency but potentially obscure structural biases. In this paper, we propose EquiLense, a post-hoc, model-agnostic framework that bridges these perspectives through clinical similarity matching and comprehensive fairness auditing. Our method introduces the Mean Predicted Probability Difference (MPPD), which quantifies prediction inconsistencies between clinically similar patients across demographic groups, integrating both individual-level consistency and group-level equity assessment. Moreover, we provide flexible similarity matching using clinical features and comprehensive visualization tools that support practical deployment in healthcare settings. Applied to electronic health record data from over 59,000 surgical patients, our framework revealed disparities in prediction consistency even when overall model performance appeared strong. EquiLense identified differences in predicted probabilities between clinically similar patients from different racial groups, disparities that were substantially reduced when sensitive attributes were excluded from model training. Our method provides a clinically relevant and interpretable approach to fairness auditing that enables healthcare practitioners to identify, understand, and address algorithmic disparities in real-world deployment settings.
Goretsky, A.; Dmitrienko, A.; Tang, I.; Lari, N.; Kunhardt, O.; Rashid Khan, R.; Marcussen, C. E.; Catto, A.; Mallia, D.; Leshchenko, A.; Lin, A.; Raja, A.; Salleb-Aouissi, A.; Pe'er, I.; Wapner, R.; Gyamfi-Bannerman, C.
Show abstract
In 2010, the Eunice Kennedy Shriver National Institute of Child Health and Human Development (NICHD) started the Nulliparous Pregnancy Outcomes Study: Monitoring Mothers-to-be (nuMoM2b), a prospective cohort study of a racially/ethnically/geographically diverse population of nulliparous women with singleton gestation. The nuMoM2b is a very large dataset, consisting of data for 10,038 patients with over 4,600 features per patient, spread out over 80 files. In this report, we share our experience preparing and working with this dataset. We present our data preprocessing of the nuMoM2b dataset to get a deeper understanding of the data, extract the most relevant features, make the fewest assumptions when filling in unknown values, and reducing the dimensionality of the data. We hope this report is useful to researchers interested in building machine learning and statistical models from the nuMoM2b dataset.
Boyce, D.; Premasiri, A.; Sullivan, S.; Levine, B.; Vieira, F. G.
Show abstract
Objectives: Patient-directed SMART on FHIR lets registries acquire longitudinal electronic health record data, but the payload requires substantial engineering before use. We present Registry Forge, an open-source pipeline that converts it into research-ready outputs. Materials and Methods: Registry Forge decodes and parses mixed C-CDA, HTML, RTF, PDF, and FHIR inputs, joins records to a canonical patient identifier, and emits a browser-viewable dashboard, an OMOP CDM v5.4 data set, GA4GH Phenopackets v2, a code inventory, and regex extractions of disease-specific narrative content. Results: Applied to the ALS Research Collaborative Study (94 participants, 56 US health systems), it processed 22,686 source files and 1,791 FHIR Bundles (109,599 resources); only 15.0% of files were full C-CDA. Discussion: This pipeline generalizes to any registry acquiring data through patient-directed SMART on FHIR. Conclusion: Registry Forge closes the acquisition-to-analysis gap with no server infrastructure and is openly available.
Chatham, A. H.; Bradley, E. D.; Schirle, L.; Sanchez-Roige, S.; Samuels, D. C.; Jeffery, A. D.
Show abstract
ImportanceIndividuals whose chronic pain is managed with opioids are at high risk of developing an opioid use disorder. Large data sets, such as electronic health records, are required for conducting studies that assist with identification and management of problematic opioid use. ObjectiveDetermine whether regular expressions, a highly interpretable natural language processing technique, could automate a validated clinical tool (Addiction Behaviors Checklist1) to expedite the identification of problematic opioid use in the electronic health record. DesignThis cross-sectional study reports on a retrospective cohort with data analyzed from 2021 through 2023. The approach was evaluated against a blinded, manually reviewed holdout test set of 100 patients. SettingThe study used data from Vanderbilt University Medical Centers Synthetic Derivative, a de-identified version of the electronic health record for research purposes. ParticipantsThis cohort comprised 8,063 individuals with chronic pain. Chronic pain was defined by International Classification of Disease codes occurring on at least two different days.18 We collected demographic, billing code, and free-text notes from patients electronic health records. Main Outcomes and MeasuresThe primary outcome was the evaluation of the automated method in identifying patients demonstrating problematic opioid use and its comparison to opioid use disorder diagnostic codes. We evaluated the methods with F1 scores and areas under the curve - indicators of sensitivity, specificity, and positive and negative predictive value. ResultsThe cohort comprised 8,063 individuals with chronic pain (mean [SD] age at earliest chronic pain diagnosis, 56.2 [16.3] years; 5081 [63.0%] females; 2982 [37.0%] male patients; 76 [1.0%] Asian, 1336 [16.6%] Black, 56 [1.0%] other, 30 [0.4%] unknown race patients, and 6499 [80.6%] White; 135 [1.7%] Hispanic/Latino, 7898 [98.0%] Non-Hispanic/Latino, and 30 [0.4%] unknown ethnicity patients). The automated approach identified individuals with problematic opioid use that were missed by diagnostic codes and outperformed diagnostic codes in F1 scores (0.74 vs. 0.08) and areas under the curve (0.82 vs 0.52). Conclusions and RelevanceThis automated data extraction technique can facilitate earlier identification of people at-risk for, and suffering from, problematic opioid use, and create new opportunities for studying long-term sequelae of opioid pain management. Key PointsO_ST_ABSQuestionC_ST_ABSCan an interpretable natural language processing method automate a valid, reliable clinical tool in order to expedite the identification of problematic opioid use in the electronic health record? FindingsIn this cross-sectional study of patients with chronic pain, an automated natural language processing approach identified individuals with problematic opioid use that were missed by diagnostic codes. MeaningRegular expressions can be used in automatically identifying problematic opioid use in an interpretable and generalizable manner.
Brandt, P. S.; Kho, A. N.; Luo, Y.; Pacheco, J. A.; Walunas, T. L.; Hakonarson, H.; Hripcsak, G.; Liu, C.; Shang, N.; Weng, C.; Walton, N.; Carrell, D. S.; Crane, P. K.; Larson, E.; Chute, C. G.; Kullo, I.; Carroll, R.; Denny, J. C.; Ramirez, A.; Wei, W.-Q.; Pathak, J.; Wiley, L. K.; Richesson, R.; Starren, J. B.; Rasmussen, L. V.
Show abstract
ObjectiveAnalyze a publicly available sample of rule-based phenotype definitions to characterize and evaluate the types of logical constructs used. Materials & MethodsA sample of 33 phenotype definitions used in research and published to the Phenotype KnowledgeBase (PheKB), that are represented using Fast Healthcare Interoperability Resources (FHIR) and Clinical Quality Language (CQL) was analyzed using automated analysis of the computable representation of the CQL libraries. ResultsMost of the phenotype definitions include narrative descriptions and flowcharts, while few provide pseudocode or executable artifacts. Most use 4 or fewer medical terminologies. The number of codes used ranges from 5 to 6865, and value sets from 1 to 19. We found the most common expressions used were literal, data, and logical expressions. Aggregate and arithmetic expressions are the least common. Expression depth ranges from 4 to 27. DiscussionDespite the range of conditions, we found that all of the phenotype definitions consisted of logical criteria, representing both clinical and operational logic, and tabular data, consisting of codes from standard terminologies and keywords for natural language processing. The total number and variety of expressions is low, which may be to simplify implementation, or authors may limit complexity due to data availability constraints. ConclusionThe phenotypes analyzed show significant variation in specific logical, arithmetic and other operators, but are all composed of the same high-level components, namely tabular data and logical expressions. A standard representation for phenotype definitions should support these formats and be modular to support localization and shared logic.
Murray, M. R.; Ho, Y.-L.; Panickan, V. A.; Heise, D.; Connatser, K.; Muralidhar, S.; Honerlaw, J.; Cho, K.
Show abstract
High-throughput phenotyping strategies are capable of classifying large volumes of patients. However, translating this data to real world applications is challenging. We have developed GeoPheno, a tool which displays the geospatial prevalences of EHR-based phenotypes in the Veteran population over time. Our flexible tool can display data from a wide array of phenotypes and is integrated with the CIPHER phenotype library, allowing users to view the definitions of the conditions being visualized.
Bhattacharya, B.; DeLong, G.; Mitchell, E. G.; Munia, T. T. K.; Shetty, G.; Tariq, A.
Show abstract
For the NIH Long COVID Computational Challenge (L3C) in the Fall of 2022, we developed a machine learning model to predict who is at high risk for developing Long COVID, optimized for clinical deployment. Our submission won second prize in the competition. We present lessons learned, with details on the features, model selection and performance, fairness analysis, limitations, and deployment implications.
Chhablani, C.; Shahid, U.; Parde, N.; Muslmani, S.; Hu, H.; Thorpe, D.; Afshar, M.; Karnik, N. S.; chhabra, n.
Show abstract
ObjectiveEmergency department (ED) encounters represent valuable opportunities to initiate evidence-based treatments for patients with opioid misuse, but few receive such care. Universal manual screening has been proposed to improve patient identification but is uncommon due to its time and resource-intensive nature. We sought to determine the feasibility of identifying patients with opioid misuse at the time of ED triage using machine learning (ML). MethodsWe conducted a retrospective cohort study of 1,123 ED encounters (September 2020 - March 2023) at a tertiary hospital. Encounters were enriched for opioid misuse, manually annotated, and chronologically split for training, validation, and testing. Candidate triage-time features included patient demographics, Emergency Severity Index, arrival time of day, chief complaint, comorbidities, and chronic medications. Model performance was evaluated using F1 score, area under the precision-recall curve (AUPRC), accuracy, recall, and AUROC. Post-hoc explainability analyses included SHapley Additive exPlanations (SHAP) and feature importance. ResultsAll models performed comparably to opioid-related diagnosis codes placed at any time during the encounter. Random Forest (F1=0.75 [95%CI 0.70-0.83], AUPRC=0.88 [0.81-0.93], accuracy=0.79 [0.70-0.83]) and Gradient Boosting (F1=0.77 [0.71-0.82], AUPRC=0.89 [0.85-0.93], accuracy=0.81 [0.720.84]) had among the highest F1 score and AUPRC but confidence intervals overlapped with other methods. Explainability analyses highlighted prior drug-use diagnosis codes, triage acuity, and age as top predictors. ConclusionML classifiers leveraging routinely collected triage data offer a feasible alternative to manual screening in flagging opioid misuse before physician evaluation, potentially enabling early harm-reduction interventions. Prospective multi-site validation, calibration, and bias assessments are warranted.
Wang, T. D.; Henderson, D. W.; Weber, G. M.; Morris, M.; Sadhu, E.; Murphy, S. N.; Visweswaran, S.; Klann, J. G.
Show abstract
ObjectiveFederated research networks, like Evolve to Next-Gen Accrual of patients to Clinical Trials (ENACT), aim to facilitate medical research by exchanging electronic health record (EHR) data. However, poor data quality can hinder this goal. While networks typically set guidelines and standards to address this problem, we developed an organically evolving, data-centric method using patient counts to identify data quality issues, applicable even to sites not yet in the network. Materials and MethodsWe distribute high-performance patient counting scripts as part of Integrating Biology at the Bedside (i2b2), which all ENACT sites operate. They produce counts of patients associated with ENACT ontology terms for each site. At the ENACT Hub, our pipeline aggregates site-contributed counts to produce network statistics, which our self-service web application, Data Quality Explorer (DQE), ingests to help sites conduct data quality investigation relative to the network. ResultsThirteen ENACT sites have contributed their patient counts, and currently seven sites have signed up to use DQE to analyze data quality issues. We announced a call to all ENACT sites to contribute additional patient counts. DiscussionIdentifying site data quality problems relative to the network is novel. Using a metric based on evolving network statistics complements rigid data quality checks. It is adaptable to any network and has low barriers of entry, with patient counting being the sole requirement. ConclusionWe implemented a metric for conducting data quality investigation in ENACT using patient counting and network statistics. Our end-to-end pipeline is privacy-preserving and the underlying design is generalizable.